[1] 0.1731224
[1] 0.001043532
University of Amsterdam
11 September 2026
In this lecture we discuss:
Reading: Chapter 7 (§7.1–7.8), 5.7 (Scatterplots), 2.8–2.9 (Standard error & confidence intervals)
Not exam material: §7.2.5, §7.4.3–7.4.5 (Spearman, Kendall, point-biserial), §7.6 (comparing correlations).
[1] 0.1731224
[1] 0.001043532
| \(r\) | \(p\) | |
|---|---|---|
| 7 damaged flights | -0.58 | .173 |
| All 23 flights | -0.64 | .001 |
Data: challenger.csv, from the Presidential Commission on the Space Shuttle Challenger Accident (1986, Vol. 1: 129-131), via Dalal, Fowlkes & Hoadley (1989).
What is the difference between

In probability and statistics, Student’s t-distribution (or simply the t-distribution) is any member of a family of continuous probability distributions that arises when estimating the mean of a normally distributed population in situations where the sample size is small and population standard deviation is unknown.
In the English-language literature it takes its name from William Sealy Gosset’s 1908 paper in Biometrika under the pseudonym “Student”. Gosset worked at the Guinness Brewery in Dublin, Ireland, and was interested in the problems of small samples, for example the chemical properties of barley where sample sizes might be as low as 3.
Source: Wikipedia
set.seed(123)
layout(matrix(c(2:6,1,1,7:8,1,1,9:13), 4, 4))
n <- 50 # Sample size
df <- n - 1 # Degrees of freedom
mu <- 100
sigma <- 15
IQ <- seq(mu-45, mu+45, 1)
par(mar=c(4,2,2,0))
plot(IQ, dnorm(IQ, mean = mu, sd = sigma), type='l', col="black", main = "Population Distribution",
lwd = 3, bty= "n", axes = FALSE, xlab = c(60, 140))
axis(1)
n.samples <- 12
for(i in 1:n.samples) {
par(mar=c(2,2,2,0))
hist(rnorm(n, mu, sigma), main="Sample Distribution", xlab = c(60, 140),
cex.axis=.5, col=rainbow(12)[i], cex.main = .75, las = 1, axes = FALSE)
axis(1)
}
Let’s take one sample from our normal population.
[1] 116.11018 99.58980 99.50004 77.25899 111.85578 96.83899 90.14886
[8] 78.81961 95.50356 87.26408 94.04454 81.73600 125.31384 99.75996
[15] 116.12418 60.97450 93.20203 89.86777 81.65611 123.19914 78.77077
[22] 104.77585 112.69654 102.67285 86.87117 114.11749 102.55882 84.04753
[29] 79.17926 131.30076 89.82245 72.16643 107.99889 104.65345 79.69248
[36] 70.85565 98.25546 117.09094 109.54186 92.60594 87.48718 104.06600
[43] 102.36030 109.44568 94.06303 113.49031 87.53783 95.04183 111.11222
[50] 114.84957

More samples, keeping \(\bar{x}\) each time.
The means have their own distribution.

That spread has a name.
95% confidence interval
\[SE = \frac{\text{Standard deviation}}{\text{Square root of sample size}} = \frac{s}{\sqrt{n}}\]

So if the population is normaly distributed (assumption of normality) the t-distribution represents the deviation of sample means from the population mean (\(\mu\)), given a certain sample size (\(df = n - 1\)).
The t-distibution therefore is different for different sample sizes and converges to a standard normal distribution if sample size is large enough.
The t-distribution is defined by:
\[\textstyle\frac{\Gamma \left(\frac{\nu+1}{2} \right)} {\sqrt{\nu\pi}\,\Gamma \left(\frac{\nu}{2} \right)} \left(1+\frac{x^2}{\nu} \right)^{-\frac{\nu+1}{2}}\!\]
where \(\nu\) is the number of degrees of freedom and \(\Gamma\) is the gamma function.
Source: wikipedia
Fatter tails for small \(n\) (few degrees of freedom); converges to the standard normal as \(n\) grows.
Same logic as the binomial, now continuous: stricter \(\alpha\), region further into the tails.
The same reasoning as the binomial applet, now for a continuous outcome.
Play around with this app to see how the t-distribution, \(\alpha\), and power hang together

In statistics, the Pearson correlation coefficient, also referred to as the Pearson’s r, Pearson product-moment correlation coefficient (PPMCC) or bivariate correlation, is a measure of the linear correlation between two variables X and Y. It has a value between +1 and −1, where 1 is total positive linear correlation, 0 is no linear correlation, and −1 is total negative linear correlation. It is widely used in the sciences. It was developed by Karl Pearson from a related idea introduced by Francis Galton in the 1880s.
Source: Wikipedia
\[r_{xy} = \frac{{COV}_{xy}}{S_xS_y}\] Where \(S\) is the standard deviation and \(COV\) is the covariance.
\[{COV}_{xy} = \frac{\sum_{i=1}^N (x_i - \bar{x})(y_i - \bar{y})}{N-1}\]
\[(x_i - \bar{x})(y_i - \bar{y})\]
\[{COV}_{xy} = \frac{\sum_{i=1}^N (x_i - \bar{x})(y_i - \bar{y})}{N-1}\]
\[r_{xy} = \frac{{COV}_{xy}}{S_xS_y}\]
\[r_{xy} = \frac{{COV}_{xy}}{S_xS_y}\] \[{COV}_{xy} = \frac{\sum_{i=1}^N (x_i - \bar{x})(y_i - \bar{y})}{N-1}\]
[1] -0.4409934
[1] -0.4409934
A test statistic with a known probability distribution (the t-distribution).
It’s a standardized measure of how closely our observed statistic matches the value claimed by \(H_0\):
\[t = \frac{\text{estimate} - H_0 \text{ value}}{\text{SE}(\text{estimate})}\]
We convert \(r\) to a \(t\)-statistic because \(t\) has a sampling distribution: the \(t\)-distribution with df = N-2:
\[t_r = \frac{r \sqrt{N-2}}{\sqrt{1 - r^2}}\]
\[{df} = N - 2\]
\[ \begin{aligned} H_0 &: t_r = 0 \\ H_A &: t_r \neq 0 \\ H_A &: t_r > 0 \\ H_A &: t_r < 0 \\ \end{aligned} \]
The hypotheses above translate into where we shade the rejection region:
Same \(\alpha\), split over two tails or spent on one.
Locate in \(t\)-distribution
\[P(|t| \geq 4.94 \mid H_0) < .001\]
| Flights | r | N | df | t | p |
|---|---|---|---|---|---|
| 7 damaged flights | -0.58 | 7 | 5 | -1.59 | 0.173 |
| All 23 flights | -0.64 | 23 | 21 | -3.80 | 0.001 |
Nearly the same \(r\); but \(\sqrt{N-2}\) causes the t-statistics to differ substantially.
All three are related:
How much of the association is unique to anxiety and performance?

\[\LARGE{r_{xy \cdot z} = \frac{r_{xy} - r_{xz} r_{yz}}{\sqrt{(1 - r_{xz}^2)(1 - r_{yz}^2)}}}\]
Or do anxious students simply revise less?
numerator <- cor.exam.anxiety - (cor.exam.revise * cor.anxiety.revise)
denominator <- sqrt( (1-cor.exam.revise^2)*(1-cor.anxiety.revise^2) )
partial.correlation <- numerator / denominator
partial.correlation[1] -0.2466658
Holding revision constant, the link shrinks.
Locate in t-distribution
pets.jasp can be downloaded from the JASP data libraryChamorro-Premuzic et al. (2008): Why do you like your lecturers?
Do students like lecturers who resemble themselves?
stu_*: the student’s own five personality traits (NEO-FFI)lec_*: how much they want each trait in a lecturer, from −5 to +5
Scientific & Statistical Reasoning